Questionnaire discrimination : ( re ) - introducing coefficient delta
نویسنده
چکیده
Background: Questionnaires are used routinely in clinical research to measure health status and quality of life. Questionnaire measurements are traditionally formally assessed by indices of reliability (the degree of measurement error) and validity (the extent to which the questionnaire measures what it is supposed to measure). Neither of these indices assesses the degree to which the questionnaire is able to discriminate between individuals, an important aspect of measurement. This paper introduces and extends an existing index of a questionnaire's ability to distinguish between individuals, that is, the questionnaire's discrimination. Methods: Ferguson (1949) [1] derived an index of test discrimination, coefficient δ, for psychometric tests with dichotomous (correct/incorrect) items. In this paper a general form of the formula, δG, is derived for the more general class of questionnaires allowing for several response choices. The calculation and characteristics of δG are then demonstrated using questionnaire data (GHQ-12) from 2003–2004 British Household Panel Survey (N = 14761). Coefficients for reliability (α) and discrimination (δG) are computed for two commonly-used GHQ-12 coding methods: dichotomous coding and four-point Likert-type coding. Results: Both scoring methods were reliable (α > 0.88). However, δG was substantially lower (0.73) for the dichotomous coding of the GHQ-12 than for the Likert-type method (δG = 0.96), indicating that the dichotomous coding, although reliable, failed to discriminate between individuals. Conclusion: Coefficient δG was shown to have decisive utility in distinguishing between the crosssectional discrimination of two equally reliable scoring methods. Ferguson's δ has been neglected in discussions of questionnaire design and performance, perhaps because it has not been implemented in software and was restricted to questionnaires with dichotomous items, which are rare in health care research. It is suggested that the more general formula introduced here is reported as δG, to avoid the implication that items are dichotomously coded. Background Questionnaire measures are routinely used in clinical research as measures of health status and quality of life [2] as well as other outcomes such as mood, stress, satisfaction and so on. The theory underlying the use of questionnaires as instruments of measurement is predominantly psychometric [3], and in keeping with this tradition the measurement properties of such questionnaires are Published: 18 May 2007 BMC Medical Research Methodology 2007, 7:19 doi:10.1186/1471-2288-7-19 Received: 15 January 2007 Accepted: 18 May 2007 This article is available from: http://www.biomedcentral.com/1471-2288/7/19 © 2007 Hankins; licensee BioMed Central Ltd. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. BMC Medical Research Methodology 2007, 7:19 http://www.biomedcentral.com/1471-2288/7/19 Page 2 of 5 (page number not for citation purposes) reported as indices of reliability and validity. The reliability coefficient (for example, Cronbach's α) estimates the degree of measurement error in the data, and hence the reproducibility of the measurements. Validity refers to the degree to which the questionnaire measures what is intended to be measured, and this is usually inferred from the degree to which the questionnaire agrees with other criteria. Reliability and validity of measurement are of course paramount for good-quality data, but the degree to which a measurement instrument is capable of discerning differences between individuals is also a fundamental aspect of measurement theory [4]. For a questionnaire to be useful in assessing health status, it must be able to distinguish between individuals who differ in health status, and fail to distinguish between those who do not. A questionnaire that failed to distinguish real differences would be unlikely to be valid, and hence discrimination is a necessary but not sufficient condition of validity. The concept described here as 'discrimination' is also referred to as 'discriminatory power' [3] but should not be confused with discriminant validity, item discrimination or discriminant functions. A little-reported statistic, Ferguson's [1]δ, quantifies the extent to which a measure can distinguish between cases. The statistic is conceptually simple. It is the ratio of observed differences to the theoretical maximum possible number of differences. When all possible scores occur with the same frequency, then the scale is maximally discriminating and the index is 1.0. Ferguson demonstrated that a normal distribution of test scores would yield a coefficient of around 0.9, and a rectangular distribution, 1.0. Skewed distributions result in fewer discriminations and hence lower values of δ, reaching a minimum of 0.0 when no discriminations at all are made and every respondent has the same score. That this statistic has not been more widely used may be due to the limiting assumption that the measure comprises dichotomous items (e.g. incorrect/correct), with each response coded as 0 or 1. Most health status questionnaires use polytomous scales, typically fiveor sevenpoint Likert-type scales (e.g. Strongly disagree, Disagree, Not sure, Agree, Strongly Agree). Researchers wishing to compute δ would therefore be forced to dichotomise item responses in order to compute the statistic. As noted above, discrimination does not ensure validity: a high δ indicates that something is being discriminated, but not necessarily the thing intended. As Guilford [5] points out, any discussion of discrimination must take place within the more problematic context of validity. Interestingly, Guilford also suggested that the goals of maximising both discrimination and reliability may be incompatible. High reliability is sometimes claimed when the measure is constructed of highly-correlated items. As well as potentially limiting the validity of the resulting scale by excluding uncorrelated but valid items, this will tend to decrease discrimination. Depending on the circumstances it may be desirable to improve discrimination by increasing the heterogeneity of the questionnaire items at the cost of reliability (although reliability should not fall below an acceptable level). Hence discrimination should be a key consideration of questionnaires at the design stage. The remainder of this paper develops the original formula for δ to allow for the computation of the statistic for questionnaire measures with polytomous items. The resulting general formula applies equally well to dichotomous and polytomous scales. The utility of the statistic will then be demonstrated using data from the 12-item General Health Questionnaire (GHQ-12) [6], which may be coded as the sum of 12 dichotomous items (known as 0011 coding) or of 12 items with four response categories (known as 0123 coding). Methods Ferguson's formula for δ assumes that the test comprises one or more items, each with only two response categories: incorrect or correct. The items are therefore dichotomous and coded as 0 or 1, respectively. The definitional formula for δ is: In which: n = sample size f = frequency of score i k = number of questionnaire items This definitional formula has been further modified [5,7]. Guilford simplifies it to a computational formula as follows [5]: The simplification offered by Cliff [7] is not presented here due to notational differences between his paper and δ = −
منابع مشابه
Questionnaire discrimination: (re)-introducing coefficient δ
BACKGROUND Questionnaires are used routinely in clinical research to measure health status and quality of life. Questionnaire measurements are traditionally formally assessed by indices of reliability (the degree of measurement error) and validity (the extent to which the questionnaire measures what it is supposed to measure). Neither of these indices assesses the degree to which the questionna...
متن کاملDiscrimination and reliability: equal partners? Understanding the role of discriminative instruments in HRQoL research: can Ferguson's Delta help? A response
A response to Norman GR 'Discrimination and reliability: equal partners?' and Wyrwich KW 'Understanding the role of discriminative instruments in HRQoL research: can Ferguson's Delta help?' Response I would like to thank Norman and Wyrwich for their close reading of my article [1], and also the editors for inviting this debate. It is a welcome opportunity to clarify some points and expand upon ...
متن کاملEQ-logics with delta connective
In this paper we continue development of formal theory of a special class offuzzy logics, called EQ-logics. Unlike fuzzy logics being extensions of theMTL-logic in which the basic connective is implication, the basic connective inEQ-logics is equivalence. Therefore, a new algebra of truth values calledEQ-algebra was developed. This is a lower semilattice with top element endowed with two binary...
متن کاملValidation of Health Behavior and Stages of Change Questionnaire
BACKGROUND The transtheoretical model (TTM) has been widely used to promote healthy behaviors in different groups. However, a questionnaire has not yet been developed to evaluate the health behaviors that medical practitioners often consider in individuals with cancer or at a high risk of developing cancer. PURPOSE The aim of this study was to construct and validate the Health Behavior and St...
متن کاملReliability of a self-report Italian version of the AUDIT-C questionnaire, used to estimate alcohol consumption by pregnant women in an obstetric setting Valutazione dell’affidabilità della versione italiana del questionario AUDIT-C per la rilevazione del consumo di alcol in gravidanza
Aim. Alcohol consumption during pregnancy can result in a range of harmful effects on the developing foetus and newborn, called Fetal Alcohol Spectrum Disorders (FASD). The identification of pregnant women who use alcohol enables to provide information, support and treatment for women and the surveillance of their children. The AUDIT-C (the shortened consumption version of the Alcohol Use Disor...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2017